Search CORE

43 research outputs found

Workload-Aware Performance Tuning for Autonomous DBMSs

Author: Chainani Naresh
Lin Chunbin
Lu Jiaheng
Yan Zhengtong
Publication venue: IEEE
Publication date: 19/04/2021
Field of study

Optimal configuration is vital for a DataBase Management System (DBMS) to achieve high performance. There is no one-size-fits-all configuration that works for different workloads since each workload has varying patterns with different resource requirements. There is a relationship between configuration, workload, and system performance. If a configuration cannot adapt to the dynamic changes of a workload, there could be a significant degradation in the overall performance of DBMS unless a sophisticated administrator is continuously re-configuring the DBMS. In this tutorial, we focus on autonomous workload-aware performance tuning, which is expected to automatically and continuously tune the configuration as the workload changes. We survey three research directions, including 1) workload classification, 2) workload forecasting, and 3) workload-based tuning. While the first two topics address the issue of obtaining accurate workload information, the third one tackles the problem of how to properly use the workload information to optimize performance. We also identify research challenges and open problems, and give real-world examples about leveraging workload information for database tuning in commercial products (e.g., Amazon Redshift). We will demonstrate workload-aware performance tuning in Amazon Redshift in the presentation.Peer reviewe

Helsingin yliopiston digitaalinen arkisto

Optimal algorithms for selecting top-k combinations of attributes : theory and applications

Author: Lin Chunbin
Lu Jiaheng
Wang Jianguo
Wei Zhewei
Xiao Xiaokui
Publication venue
Publication date: 01/01/2017
Field of study

Traditional top-k algorithms, e.g., TA and NRA, have been successfully applied in many areas such as information retrieval, data mining and databases. They are designed to discover k objects, e.g., top-k restaurants, with highest overall scores aggregated from different attributes, e.g., price and location. However, new emerging applications like query recommendation require providing the best combinations of attributes, instead of objects. The straightforward extension based on the existing top-k algorithms is prohibitively expensive to answer top-k combinations because they need to enumerate all the possible combinations, which is exponential to the number of attributes. In this article, we formalize a novel type of top-k query, called top-k, m, which aims to find top-k combinations of attributes based on the overall scores of the top-m objects within each combination, where m is the number of objects forming a combination. We propose a family of efficient top-k, m algorithms with different data access methods, i.e., sorted accesses and random accesses and different query certainties, i.e., exact query processing and approximate query processing. Theoretically, we prove that our algorithms are instance optimal and analyze the bound of the depth of accesses. We further develop optimizations for efficient query evaluation to reduce the computational and the memory costs and the number of accesses. We provide a case study on the real applications of top-k, m queries for an online biomedical search engine. Finally, we perform comprehensive experiments to demonstrate the scalability and efficiency of top-k, m algorithms on multiple real-life datasets.Peer reviewe

Helsingin yliopiston digitaalinen arkisto

DR-NTU (Digital Repository of NTU)

Synergy of Database Techniques and Machine Learning Models for String Similarity Search and Join

Author: Li Chen
Lin Chunbin
Lu Jiaheng
Wang Jin
Publication venue: ACM
Publication date: 01/11/2019
Field of study

Peer reviewe

Crossref

Helsingin yliopiston digitaalinen arkisto